Papers by Jey Han Lau

45 papers
Controlling Distributional Bias in Multi-Round LLM Generation via KL-Optimized Fine-Tuning (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods focus on single-round inference, but this view is problematic in real-world applications.
Approach: They propose a framework that couples Steering Token Calibration with Semantic Alignment to ensure that LLMs are correctly aligned across gender, race, and sentiment.
Outcome: The proposed framework outperforms baseline methods in achieving precise distributional control in attribute generation tasks.
An Interpretable Neuro-Symbolic Reasoning Framework for Task-Oriented Dialogue Generation (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to interpret task-oriented dialogue systems employ an implicit reasoning strategy that makes the model predictions uninterpretable to humans.
Approach: They propose a neuro-symbolic approach that performs explicit reasoning that justifies model decisions by reasoning chains.
Outcome: The proposed approach achieves better results and introduces an interpretable decision process.
The Influence of Context on Sentence Acceptability Judgements (P18-2)

Copied to clipboard

Challenge: a paper examining the influence of document context on acceptability judgements for English sentences is published in journal journal of linguistics.
Approach: They propose to use document context to assess acceptability judgements for English sentences . they also test the accuracy of neural models that incorporate document context during training .
Outcome: The proposed model improves acceptability ratings for ill-formed sentences, but reduces them for well-formed ones.
LipKey: A Large-Scale News Dataset for Absent Keyphrases Generation and Abstractive Summarization (2022.coling-1)

Copied to clipboard

Challenge: Existing work has addressed each element individually, but this study focuses on LipKey, the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles.
Approach: They propose a novel news dataset that consists of highly absent keyphrases . they combine lips keyphrase and TF-IDF to obtain abstractive summaries .
Outcome: The proposed dataset is the largest news corpus with human-written abstractive summaries, absent keyphrases, and titles.
COMMUNITYNOTES: A Dataset for Exploring the Helpfulness of Fact-Checking Explanations (2026.findings-eacl)

Copied to clipboard

Challenge: X, Meta, and TikTok are experimenting with community-based factchecking . community-driven verification is a way to provide explanatory notes that clarify why a post might be misleading .
Approach: They propose a framework that optimizes the helpfulness of explanatory notes and the reason for this by automatically optimizing the prompt definitions.
Outcome: The proposed framework improves helpfulness and reason prediction on 104k posts with user-provided notes and helpfulness labels.
Liputan6: A Large-scale Indonesian Dataset for Text Summarization (2020.aacl-main)

Copied to clipboard

Challenge: Despite having the fourth largest speaker population in the world, 1 Indonesian is under-represented in NLP.
Approach: They propose to use a large-scale Indonesian summarization dataset to test extractive and abstractive summarizing methods.
Outcome: The proposed methods are compared with multilingual and monolingual BERT-based models.
M3: Multi-level dataset for Multi-document summarisation of Medical studies (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing summarisation systems are not up to such complex tasks, yet limited tools exist to determine where and why they are failing.
Approach: They propose to use a dataset to evaluate the quality of summarisation systems in the biomedical domain.
Outcome: The proposed model can be used to evaluate the quality of summarisation systems in the biomedical domain.
DUCK: Rumour Detection on Social Media by Modelling User and Comment Propagation Networks (2022.naacl-main)

Copied to clipboard

Challenge: Social media rumours can cause significant economic and social disruption.
Approach: They propose a rumour detection algorithm that leverages transformers and graph attention networks to jointly model social media conversations and the network of users who engaged in them.
Outcome: The proposed algorithm produces superior performance over four widely used benchmark rumour datasets in English and Chinese.
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages (2023.eacl-main)

Copied to clipboard

Challenge: In Indonesia, many languages are endangered and some are even extinct due to the unavailability of data resources and benchmarks.
Approach: They propose a high-quality multilingual parallel corpus that covers 10 local languages from Indonesia.
Outcome: The proposed resource includes sentiment and machine translation datasets, and bilingual lexicons.
A Sentiment Consolidation Framework for Meta-Review Generation (2024.acl-long)

Copied to clipboard

Challenge: Recent advances in abstractive text summarization have created plausible summaries, but it is unclear if they truly possess the capability of information consolidation to generate summary.
Approach: They propose to prompt large language models to generate meta-reviews and use evaluation metrics to assess the quality of generated meta- reviews.
Outcome: The proposed framework proves that human meta-reviewers follow a framework of sentiment consolidation to write meta- reviews compared with prompting them with simple instructions.
Factual Dialogue Summarization via Learning from Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing models generate fluent and coherent summaries, but inconsistencies can be found in generated summary.
Approach: They propose to use symbolic knowledge distillation to improve the factual consistency of smaller pretrained models for dialogue summarization.
Outcome: The proposed model outperforms baseline models in BART, PEGASUS, and Flan-T5 in factual consistency and accuracy.
One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia (2022.acl-long)

Copied to clipboard

Challenge: There are more than 700 languages spoken in Indonesia, equal to 10% of the world's languages, second only to Papua New Guinea.
Approach: They focus on the languages spoken in Indonesia, the world's second most linguistically diverse nation, and the fourth most populous nation of the world.
Outcome: The proposed model is based on the languages spoken in Indonesia, the world's second-most linguistically diverse nation, with 273 million people spread over 17,508 islands.
On the Interplay between Human Label Variation and Model Fairness (2026.findings-eacl)

Copied to clipboard

Challenge: Existing studies on the impact of human label variation on model fairness have not explored the interaction between HLV and performance.
Approach: They compare human label variation (HLV) training methods with four other methods . they find that HLV methods improve performance without harming fairness .
Outcome: The proposed methods improve fairness without explicit debiasing under certain configurations.
The patient is more dead than alive: exploring the current state of the multi-document summarisation of the biomedical literature (2022.acl-long)

Copied to clipboard

Challenge: Existing evaluation approaches to multi-document summarization of biomedical literature lack consistency and transparency.
Approach: They propose a systematic approach to human evaluation of biomedical summaries and apply it to analyze the summary generated by two current evaluation models.
Outcome: The proposed evaluation framework is based on two state-of-the-art models and examines the summaries generated by the two models to understand the deficiencies of existing evaluation approaches.
Unsupervised Paraphrasing of Multiword Expressions (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for paraphrasing multiword expressions in context are unsupervised . multiwords are notoriously difficult to model because the meaning of the whole can diverge substantially from that of the component words.
Approach: They propose an unsupervised approach to paraphrasing multiword expressions in context using monolingual corpus data and pre-trained language models.
Outcome: The proposed method outperforms all unsupervised systems and rivals supervised systems on the SemEval 2022 idiomatic text similarity task.
Not another Negation Benchmark: The NaN-NLI Test Suite for Sub-clausal Negation (2022.aacl-main)

Copied to clipboard

Challenge: Negation is an important linguistic phenomenon which denotes non-existence, denial, or contradiction.
Approach: They propose a natural language inference test suite to test models for negation . they use a linguistic framework to analyze negation types and constructions .
Outcome: The proposed test suite is more challenging than existing benchmarks on negation . it includes annotation of negation types and constructions grounded in linguistic theory .
WET: Overcoming Paraphrasing Vulnerabilities in Embeddings-as-a-Service with Linear Transformation Watermarks (2025.acl-long)

Copied to clipboard

Challenge: Existing EaaS watermarks can be removed by paraphrasing when attackers clone the model.
Approach: They propose a method that integrates a target embedding into the original embeddable based on the presence of trigger words in the input text.
Outcome: The proposed technique is empirically and theoretically robust against paraphrasing.
CIG: Measuring Conversational Information Gain in Deliberative Dialogues with Semantic Memory Dynamics (2026.acl-long)

Copied to clipboard

Challenge: Using a semantic memory, we score each utterance along three interpretable dimensions: Novelty, Relevance, and Implication Scope.
Approach: They propose a framework for Conversational Information Gain that evaluates each utterance in terms of how it advances collective understanding of the target topic.
Outcome: The proposed framework evaluates each utterance in terms of how it advances collective understanding of the target topic.
Grey-box Adversarial Attack And Defence For Sentiment Classification (2021.naacl-main)

Copied to clipboard

Challenge: Recent advances in deep neural networks have created applications for a range of different domains.
Approach: They propose a grey-box adversarial attack and defence framework for sentiment classification . they show that the framework produces an improved classifier that is robust in defending .
Outcome: The proposed framework produces an improved classifier that is robust in defending against multiple adversarial attacking methods.
Early Rumour Detection (N19-1)

Copied to clipboard

Challenge: Existing studies on rumour detection are concerned with timing, but few are interested in how early we can detect them.
Approach: They propose a method that integrates reinforcement learning to learn the minimum number of posts required before classifying an event as a rumour.
Outcome: The proposed model detects rumours earlier than state-of-the-art systems while maintaining comparable accuracy.
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task (2025.findings-acl)

Copied to clipboard

Challenge: Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process.
Approach: They propose a multi-stage AVR benchmark based on RAVEN to assess reasoning across varying levels of complexity.
Outcome: The proposed metric considers the correctness of intermediate steps in addition to the final outcomes.
Unsupervised Lexical Substitution with Decontextualised Embeddings (2022.coling-1)

Copied to clipboard

Challenge: Existing methods for lexical substitution using pre-trained language models have some limitations.
Approach: They propose an unsupervised method for lexical substitution using pre-trained language models.
Outcome: The proposed method outperforms baseline models and establishes a state-of-the-art without supervision or fine-tuning.
Can LLMs Simulate L2-English Dialogue? An Information-Theoretic Analysis of L1-Dependent Biases (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) can simulate non-native-like English use observed in human second language (L2) learners interfered with by their native first language (N1) knowledge.
Approach: They use large language models to simulate non-native-like English use observed in human second language (L2) learners, and then compare their outputs to real L2 learner data.
Outcome: The proposed models replicate L1-dependent patterns observed in human second language (L2) learners, with distinct influences from various languages.
Top-down Discourse Parsing via Sequence Labelling (2021.eacl-main)

Copied to clipboard

Challenge: Discourse analysis is a systematic way to understand how texts are segmented hierarchically into discourse units.
Approach: They propose a top-down approach to discourse parsing that is conceptually simpler than its predecessors.
Outcome: The proposed model eliminates the decoder and reduces the search space for splitting points.
Beyond Seen Data: Improving KBQA Generalization Through Schema-Guided Logical Form Generation (2025.emnlp-main)

Copied to clipboard

Challenge: Knowledge base question answering (KBQA) aims to answer user questions in natural language using rich human knowledge stored in large KBs.
Approach: They propose a model that injects schema contexts into entity retrieval and logical form generation to enhance generalizability.
Outcome: The proposed model outperforms state-of-the-art models on two commonly used benchmark datasets across a variety of test settings.
FLUKE: A Linguistically-Driven and Task-Agnostic Framework for Robustness Evaluation (2026.findings-eacl)

Copied to clipboard

Challenge: FLUKE introduces controlled variations across linguistic levels and leverages large language models with human validation to generate modifications.
Approach: They propose a framework for assessing model robustness through systematic minimal variations of test data.
Outcome: The proposed framework evaluates models and LLMs across six diverse NLP tasks and shows that they are more robust to natural, fluent modifications than base models.
Automatic Classification of Neutralization Techniques in the Narrative of Climate Change Scepticism (2021.naacl-main)

Copied to clipboard

Challenge: neutralisation is used to justify lack of action or promote an alternative view of climate change . action on climate change has become an increasingly partisan issue with strong opposition voices discrediting scientists and spreading scepticism and misinformation.
Approach: They propose to use neutralisation techniques to introduce the problem to the nlp community and to collect manual annotations of neutralised techniques in text relating to climate change.
Outcome: The proposed models are supervised and semi-supervised by a team of researchers from the nlp and the ccsc.
IndoLEM and IndoBERT: A Benchmark Dataset and Pre-trained Language Model for Indonesian NLP (2020.coling-main)

Copied to clipboard

Challenge: despite being spoken by 200 million people, the Indonesian language is underrepresented in NLP research.
Approach: They propose a dataset for Indonesian that includes seven NLP tasks . they also propose 'indonesian language evaluation Montage' tasks that are based on previous work .
Outcome: The proposed dataset shows that IndoBERT outperforms IndoLEM over most of the tasks.
Interaction Matters: An Evaluation Framework for Interactive Dialogue Assessment on English Second Language Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing data on ESL speakers' communication and interaction skills are lacking in the evaluation of the sophisticated features of dialogue.
Approach: They propose an evaluation framework for interactive dialogue assessment in ESL speakers.
Outcome: The proposed framework provides a means to assess ESL communication, useful for language assessment.
WHoW: A Cross-domain Approach for Analysing Conversation Moderation (2025.naacl-long)

Copied to clipboard

Challenge: Using this framework, we annotated 5,657 sentences with human judges and 15,494 sentences with GPT-4o from two domains: TV debates and radio panel discussions.
Approach: They propose an evaluation framework for analyzing the facilitation strategies of moderators across different domains/scenarios by examining their motives (Why), dialogue acts (How) and target speaker (Who).
Outcome: The framework is generalisable across domains and reveals distinct modes of moderation: debate moderators emphasise coordination and facilitate interaction through questions and instructions, panel discussion moderator prioritize information provision and actively participate in discussions.
Deep-speare: A joint neural model of poetic language, meter and rhyme (P18-1)

Copied to clipboard

Challenge: a recent surge of interest in deep learning has led to creative applications for poetry generation . a novel joint architecture captures language, rhyme and meter for sonnet modelling .
Approach: They propose a joint architecture that captures language, rhyme and meter for sonnet modelling.
Outcome: The proposed architecture captures language, rhyme and meter for sonnet modelling.
Decomposed Opinion Summarization with Verified Aspect-Aware Modules (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for summarizing opinions from large-scale online reviews are not available for crowdsourcing and are difficult to crowdsource.
Approach: They propose a domain-agnostic modular approach guided by review aspects to separate tasks of aspect identification, opinion consolidation, and meta-review synthesis to enable greater transparency and ease of inspection.
Outcome: The proposed approach generates more grounded summaries than baseline models, as verified through automated and human evaluations.
Discourse Probing of Pretrained Language Models (2021.naacl-main)

Copied to clipboard

Challenge: Existing work on probing of pretrained language models has focused on sentence-level syntactic tasks.
Approach: They introduce document-level discourse probing to evaluate the ability of pretrained LMs to capture document- level relations.
Outcome: The proposed model performs best in encoder, but only in the encoder layer.
How Furiously Can Colorless Green Ideas Sleep? Sentence Acceptability in Context (2020.tacl-1)

Copied to clipboard

Challenge: a recent study shows that context affects our perception of sentence acceptability, but few studies investigate how it affects language models.
Approach: They compare acceptability ratings of sentences judged in isolation with a relevant context and with an irrelevant context.
Outcome: The proposed model achieves state-of-the-art for unsupervised acceptability prediction.
An Interpretable and Crosslingual Method for Evaluating Second-Language Dialogues (2025.naacl-long)

Copied to clipboard

Challenge: Existing studies on second language (SL) assessment of conversational fluency and interactivity have focused on written correction or pronunciation from ASR.
Approach: They propose a framework that assesses the relationships between micro-level linguistic features and macro-level interactivity labels for Chinese-as-a-second-language dialogues.
Outcome: The proposed framework is interpretable and can be adapted to other languages for second-language dialogue evaluation.
IndoBERTweet: A Pretrained Language Model for Indonesian Twitter with Effective Domain-Specific Vocabulary Initialization (2021.emnlp-main)

Copied to clipboard

Challenge: In IndoBERTweet, a pretraining model for Indonesian Twitter is extended with domain-specific vocabulary.
Approach: They propose a pretraining model that extends a monolingual Indonesian BERT model with domain-specific vocabulary.
Outcome: The proposed model can be initialized with the average BERT subword embedding five times faster than existing methods for vocabulary adaptation.
Robust Task-Oriented Dialogue Generation with Contrastive Pre-training and Adversarial Filtering (2022.findings-emnlp)

Copied to clipboard

Challenge: Task-oriented dialogue models can learn non-transferable generalizations by using shortcuts in the data.
Approach: They propose a contrastive learning framework to encourage models to ignore cues and focus on generalisable patterns.
Outcome: The proposed framework performs exceptionally well on task-oriented dialogue datasets.
Moderation Matters: Measuring Conversational Moderation Impact in English as a Second Language Group Discussion (2025.findings-acl)

Copied to clipboard

Challenge: Existing tools for ESL assessment focus on writing skills and lack in support for dynamic spoken interactions.
Approach: They propose an approach that integrates automatic ESL dialogue assessment and a framework that categorizes moderation strategies to assess conversational engagement and moderation effectiveness.
Outcome: The proposed approach integrates automatic ESL dialogue assessment and categorizes moderation strategies.
Evaluating Evidence Attribution in Generated Fact Checking Explanations (2025.naacl-long)

Copied to clipboard

Challenge: Existing fact-checking systems struggle with attribution quality, as their generated explanations can include hallucinations.
Approach: They propose a protocol to assess attribution quality in fact-checking explanations using human annotation and automatic annotation.
Outcome: The proposed protocol can be automated, the authors show . best-performing LLMs still generate explanations that are not always accurate .
Annotating and Detecting Fine-grained Factual Errors for Dialogue Summarization (2023.acl-long)

Copied to clipboard

Challenge: Existing work on factual inconsistency in abstractive summarization addresses this problem.
Approach: They propose a dataset with fine-grained factual error annotations named DIASUMFACT and an unsupervised model named ENDERANKER.
Outcome: The proposed model performs on par with the state-of-the-art models while requiring fewer resources.
Evaluating the Efficacy of Summarization Evaluation across Languages (2021.findings-acl)

Copied to clipboard

Challenge: Using multilingual summarization evaluation methods is more reliable and interpretable than manual methods.
Approach: They propose to use multilingual BERT within BERTScore to evaluate summarization evaluation metrics . they use English datasets that are not representative of modern summarizing systems .
Outcome: The proposed methods perform well across all languages, at a level above that for English.
Improving Visual-Semantic Embedding with Adaptive Pooling and Optimization Objective (2023.eacl-main)

Copied to clipboard

Challenge: Recent VSE models combine simple pooling methods with hard triplet loss to improve performance.
Approach: They propose an adaptive pooling strategy that allows the model to learn how to aggregate features through a combination of simple pooling methods.
Outcome: The proposed strategy outperforms current state-of-the-art systems on image-to-text and text-toimage retrieval.
CMA-R: Causal Mediation Analysis for Explaining Rumour Detection (2024.findings-eacl)

Copied to clipboard

Challenge: Existing studies on explainable fake news or rumour detection by and large use attention weights as explanation, but the use of attention weighted explanations is problematic.
Approach: They propose a causal mediation analysis approach to explain the decision-making process of neural models for rumour detection on Twitter by identifying salient tweets that explain model predictions and highlighting causally impactful words in the tweets.
Outcome: The proposed approach shows strong agreement with human judgements for critical tweets determining the truthfulness of stories.
Topic Intrusion for Automatic Topic Model Evaluation (D18-1)

Copied to clipboard

Challenge: Topic coherence is increasingly being used to evaluate topic models and filter topics for end-user applications.
Approach: They propose to use topic intrusion to guess an outlier topic given a document and a few topics to automate the task.
Outcome: The proposed method improves upon the state-of-the-art method and shows it can be used as an alternative to topic perplexity evaluation.
Give Me Convenience and Give Her Death: Who Should Decide What Uses of NLP are Appropriate, and on What Basis? (2020.acl-main)

Copied to clipboard

Challenge: a paper on automatic sentencing was a source of debate at EMNLP 2019 . paper examines whether particular datasets and tasks should be off-limits for NLP research .
Approach: They propose a neural model which performs structured prediction of individual charges laid against an individual and the prison term associated with each.
Outcome: The proposed model can predict the prison term associated with a given case on a large-scale dataset of real-world Chinese court cases.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations